Data is the raw material of intelligence
Think about how a child learns to recognise a dog. Nobody writes them a manual that says "four legs, fur, barks." Instead, the child sees hundreds of dogs over several years: big ones, small ones, fluffy ones, spotted ones. Over time they build up an internal sense of what makes something a dog. Learning happens through exposure to examples.
AI works exactly the same way. You cannot just tell an AI model what a cat looks like. You have to show it thousands of photographs of cats. You cannot tell it what spam email sounds like. You have to give it tens of thousands of examples of spam and non-spam email. The examples are the data, and the data is everything.
This raises the obvious question: what exactly counts as data?
"In God we trust. All others must bring data."
W. Edwards Deming, statistician and engineerThe two main types of data
All data falls into one of two broad categories. Understanding this distinction is one of the first things you need to understand before choosing an AI approach for any problem.
Organised into rows and columns. Think of a spreadsheet or a database table. Each row is one record, each column is one attribute. It is tidy, labelled, and easy to search. Computers have been handling this kind of data for decades.
Everything that does not fit neatly into rows and columns. The meaning is embedded inside the content itself, not in any predefined schema. It is harder to process but makes up the vast majority of data in the world.
There is a third category worth knowing: semi-structured data. This sits in between. It is not a rigid table, but it does have some organisational tags or markers. JSON files and XML documents are good examples. A product listing in an online store might be semi-structured: it has defined fields like "price" and "category," but the product description is free-form text.
Estimates suggest that around 80% of all data generated in the world is unstructured. This is why deep learning has become so important. It excels at finding patterns in raw images, audio and text in ways that classical machine learning methods simply cannot match.
A structured dataset up close
Let us look at a real example. The Titanic dataset, which you will use in your first coding exercise, is a classic structured dataset. Each row represents one passenger. Each column captures one fact about them.
| PassengerId | Survived | Pclass | Name | Age | Fare |
|---|---|---|---|---|---|
| 1 | 0 | 3 | Braund, Mr. Owen | 22 | 7.25 |
| 2 | 1 | 1 | Cumings, Mrs. John | 38 | 71.28 |
| 3 | 1 | 3 | Heikkinen, Miss. Laina | 26 | 7.93 |
| 4 | 1 | 1 | Futrelle, Mrs. Jacques | 35 | 53.10 |
| 5 | 0 | 3 | Allen, Mr. William | 35 | 8.05 |
Notice a few important things. The Survived column is 0 or 1, not "yes" or "no." Computers and AI models do not naturally understand words. Everything eventually gets converted to numbers. The Name column is technically unstructured text sitting inside a structured table, and that kind of mixed reality is very common in the real world.
Features and labels: the vocabulary of ML
In machine learning, we have specific names for the different columns of a dataset. Understanding these terms will help you read papers, documentation and code.
The columns that describe the thing you are analysing. In the Titanic dataset: age, ticket class, fare paid, number of siblings on board. These are what the model uses to make its prediction.
The column you are trying to predict. In the Titanic dataset: Survived (0 or 1). In a house price dataset, it would be the actual price. This is the answer the model is learning to produce.
Each individual record in your dataset. One row equals one observation. One passenger, one house, one transaction. The more observations you have, generally, the better your model can learn.
The number of features in your dataset. A dataset with 10 columns of features is 10-dimensional. High-dimensional data (hundreds or thousands of features) presents its own challenges, which we will encounter later.
How AI sees images: everything is numbers
Here is something that surprises most beginners. When you look at a photograph of a dog, you see a dog. When an AI model looks at the same photograph, it sees a grid of numbers.
A digital image is simply a grid of pixels. Each pixel has a colour, and that colour can be represented by three numbers: how much red, how much green and how much blue (the RGB values). Each value runs from 0 to 255. That means a 100 × 100 pixel image is actually a grid of 30,000 numbers (100 × 100 pixels × 3 colour channels).
The model does not see a picture. It sees a matrix of numbers. Understanding this is the key to understanding how computers process images, faces, X-rays, satellite photos and more.
How AI sees text: tokens and numbers
Text cannot be fed directly into an AI model either. Before a model can process the sentence "The cat sat on the mat," it needs to be converted into numbers. This process is called tokenisation.
A tokeniser breaks text into chunks called tokens, usually words or parts of words. Each unique token gets assigned a number from a vocabulary list. So "cat" might become 4,823 and "sat" might become 7,112. The sentence becomes a list of numbers, which the model can then do mathematics on.
Imagine a musician who can only read sheet music, not words. To communicate a story to them, you first have to translate every emotion and event into musical notes. The story is still there, just encoded in a language the musician can work with. AI models require exactly the same kind of translation, converting human-readable content into numbers the model can process.
Why data quality matters more than quantity
There is a common misconception that AI models simply need more data to get better. More is often better, but only if the data is good. The most important properties of a useful dataset are actually about quality, not size.
Relevant: Does the data actually contain signals that relate to the problem you are solving? Customer reviews are not useful data for predicting tomorrow's weather.
Representative: Does the dataset reflect the real world it is meant to represent? A model trained only on photos of faces with one skin tone will fail on faces with other skin tones. This is how bias enters AI systems.
Clean: Are there missing values, typos, duplicates, or impossible entries? Messy data produces unreliable models. In real AI projects, cleaning and preparing data consistently takes around 70–80% of the total time. You will experience this firsthand in Lesson 2.5.
Labelled (for supervised learning): For most machine learning tasks, someone has to label the data. This means a human tells the system which emails are spam, which photos contain cats, which loan applications defaulted. Labelling is expensive, time-consuming, and often underestimated.